Whole-spine parity for the UK: re-pinned incumbent, comparison instruments, signed register, and the E10 evidence (#686) - #747
Conversation
…_VERSION #747 records the uk-data release tag beside the byte identity (SOURCE_VERSION / source.version). A regeneration for another release must not inherit the committed tag, so the identity now carries the tag (from --release, or --version with explicit pins), the in-memory patch sets it, and the on-disk move is anchored to the SOURCE_VERSION assignment — never a global replacement of a tag literal, which also appears in unrelated pins such as the registry-parity pinned_version. The lockstep test binds the committed reference's source.version to the tool's constant once it exists. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
An armed UK national build crashed before its first stage: _source_pins stored the full Ledger provenance block under the 'ledger_facts' role, but role_pins_digest requires each role to carry exactly sha256 and size_bytes, so every armed run raised ValueError at input_pins_digest. The two halves landed independently — the strict pin contract with Logbook adoption (#666) and the Ledger role pin with the compile-parity wiring — and neither PR branch fires it alone, because only an armed run (--ledger-facts) reaches this path and the licensed runs were held. It reproduces on both #743 and #747 branches and on main. The pin now carries the feed's verified digest and byte size; the richer Ledger identity block already travels in safe_artifacts, source_vintages, and the diagnostics build block, so nothing is lost. This fix belongs upstream in #743 (the run-readiness PR whose runbook documents the armed command); it rides the build branch until then.
Calibrates a WS-E spine at national level with the #622/#623 seam, without the certified-input replay: the spine already carries the SPI and CGT derivations as declarative source stages (E7/E8 ported the same pinned sources, QRF stages, reviewed fences, and conservation receipts), so re-running the June-convention replay would discard the spine's certified derivation and substitute an equivalent one. Reuses the harness pieces that are posture-independent: sha-pinned inputs, Ledger register compilation at the declared calibration year, the frozen national doctrine, the validated H5 writer, US-format v6 diagnostics with the uk block, and a Logbook spool row. Scores against the incumbent on the same frozen register before the diagnostics are written, so the score block rides the diagnostics in a single pass (no two-pass sha dance). Assessment posture, stated in the build record: non-certified input, the declared battery is not enforced; the numeric fences that are honestly measurable here (target fit, weight ratio, ESS, zero-weight strata, admin anchors) are evaluated against the thresholds declared in uk/gates.json, and every non-evaluated entry is listed with its reason. The spine's twelve structural NaN columns (SPI-channel concepts undefined on the base-FRS channel) are zero-filled from a fixed allowlist with a receipt; any NaN outside the allowlist aborts. Adjudication (Maria, 2026-08-23): the June convention is transitional and retires with the completed swap; gate integration for the spine posture is follow-up work on the #747 lane.
5a546c1 to
8b893a1
Compare
How the spine is builtThe counterpart to the status diagram on #665 — same style, different question. That one maps how far the migration has got; this one maps what the spine is made of: which source each column comes from, and which kind of process puts it there. Nodes are coloured by mechanism, not by progress. It stops at the artifact. Calibration is deliberately outside the frame — everything here produces a file carrying design grossing weights, and the solve that replaces them is downstream. flowchart LR
subgraph P1["1 · Raw FRS — transcription"]
FRS(["<b>FRS 2024-25</b><br/>UKDS SN 9563 · 14 tables<br/><b>16,288 households</b>"])
SPINE0["<b>frs_spine</b><br/>entities · IDs · weights<br/>demographics · incomes<br/>85 columns"]
DERIV["<b>+ 6 derived stages</b><br/>employment · council tax<br/>disability · education<br/>legacy proxies · grant split"]
end
subgraph P2["2 · Seeded draws"]
BRMA(["<b>BRMA counts</b><br/>public aggregate"])
DRAWS["<b>frs_take_up · person_draws<br/>household_draws · frs_brma</b><br/>identity-keyed RNG<br/>17 columns"]
end
subgraph P3["3 · Wealth"]
WASD(["<b>WAS round 8</b><br/>UKDS SN 7215<br/>15,128-row donor"])
WEALTH["<b>was_wealth</b><br/>QRF · correlated rank draw<br/>13 columns"]
UPRATE["<b>regional_property_uprating</b><br/>rewrites property values"]
end
subgraph P4["4 · Consumption"]
LCFSD(["<b>LCFS 2023-24</b><br/>UKDS SN 9468<br/>4,202-row donor"])
ETBD(["<b>ETB 1977-2024</b><br/>UKDS SN 8856<br/>4,199-row 2023 donor"])
ANCH4(["<b>Anchors</b><br/>NEED energy · NHS by age/sex<br/>VAT + fare-index parameters"])
CONS["<b>lcfs_consumption</b><br/>regime-gated QRF<br/>+ NEED raking · 19 columns"]
ETBS["<b>etb_vat · etb_services</b><br/>QRF + NHS usage<br/>11 columns"]
end
subgraph P5["5 · Income — SPI channel"]
SPID(["<b>SPI 2022-23 PUT</b><br/>UKDS SN 9422<br/>HMRC taxpayer donor"])
ODS(["<b>HMRC published<br/>income tables</b>"])
LEAVES["<b>frs_hmrc_spine_leaves</b><br/>FRS-side derives · 6 columns"]
STACK["<b>spi_support_channel</b><br/><i>stacks 10,000 synthetic<br/>zero-weight households</i>"]
SPIINC["<b>hmrc_spi_income_spine</b><br/>QRF redraw on the channel<br/>14 columns"]
end
subgraph P6["6 · Capital gains + derived"]
CGTREF(["<b>HMRC CGT 2.1a + table 3</b><br/>Advani–Summers distribution<br/>SLC stocks · salsac anchor"])
CLONE["<b>cgt_incidence_clone</b><br/><i>duplicates every household ×2</i><br/>mass split 0.5"]
DONOR["<b>cgt_band_donors</b><br/><i>adds 270 top-tail rows</i>"]
GAINS["<b>hmrc_cgt_gains_spine</b><br/>gain amounts by size band"]
SALSAC["<b>salary_sacrifice · student_loans</b><br/>conversion depth · plan cohorts"]
end
subgraph P7["7 · Artifact"]
ART["<b>microcosm_uk_2024 spine</b><br/>52,846 hh · 61,211 bu · 113,649 p<br/>211 declared outputs<br/><b>design grossing weights</b>"]
IDENT["<b>Row identity closes</b><br/>16,288 + 10,000 = 26,288<br/>× 2 = 52,576<br/>+ 270 = <b>52,846</b>"]
end
subgraph P8["8 · What proves it"]
L0["<b>L0</b> determinism<br/>twins payload-identical"]
L1["<b>L1</b> identity receipts<br/>e4–e8, order-invariant"]
L2["<b>L2</b> whole-spine parity<br/><b>signed_parity</b> · 0 unsigned"]
end
CAL["<b>Calibration</b> · #623<br/>replaces the weights<br/><i>out of frame</i>"]
FRS --> SPINE0 --> DERIV --> DRAWS
BRMA --> DRAWS
DRAWS --> WEALTH
WASD --> WEALTH --> UPRATE --> CONS
LCFSD --> CONS
WASD -.->|fuel bridge| CONS
ANCH4 --> CONS
CONS --> ETBS
ETBD --> ETBS
ANCH4 --> ETBS
ETBS --> LEAVES --> STACK --> SPIINC
SPID --> SPIINC
ODS --> SPIINC
SPIINC --> CLONE --> DONOR --> GAINS --> SALSAC
CGTREF --> DONOR
CGTREF --> GAINS
CGTREF --> SALSAC
SALSAC --> ART --> IDENT
ART --> L0
ART --> L1
ART --> L2
L2 -.->|only after review| CAL
classDef source fill:#33415c,color:#fff
classDef mapped fill:#2f6f52,color:#fff
classDef drawn fill:#6b4c9a,color:#fff
classDef imputed fill:#b8721a,color:#fff
classDef rowop fill:#8b3a3a,color:#fff
classDef artifact fill:#1f3a5f,color:#fff
classDef out fill:#eee,color:#777,stroke-dasharray:5
class FRS,WASD,LCFSD,ETBD,SPID,ODS,BRMA,ANCH4,CGTREF source
class SPINE0,DERIV,LEAVES mapped
class DRAWS drawn
class WEALTH,UPRATE,CONS,ETBS,SPIINC,GAINS,SALSAC imputed
class STACK,CLONE,DONOR rowop
class ART,IDENT,L0,L1,L2 artifact
class CAL out
Three processes change the row count, and they are the entire difference between 16,288 and 52,846. Reading the colours is the point of the diagram. A green node that disagrees with the incumbent is a bug on one side. An orange node that disagrees is two estimators, settled against the donor survey. Getting that attribution wrong was the single largest source of false findings during this increment: Why calibration is out of frame. Every weight in this artifact is a design grossing weight straight off the FRS. The incumbent's weights are calibrated, so any weighted comparison between the two compares before-medicine to after-medicine — which is exactly why the parity instrument compares unweighted shares and record counts, why the weighted-totals surface stays dormant until there is a calibrated candidate, and why the spine's UC caseload looks half the incumbent's while actually handing the solve 9% more UC-positive benunits. |
|
Automated review pass (Claude Code) — two passes over the diff, one low-effort and one high-effort. No verification runs; findings are from reading the diff. Ordered by what I think should block the merge. Blocking — these change what the proofs mean1. 2. 3. Follow-ups — real, but not merge-blocking4. 5. 6. 7. On the spine itselfBoth passes went looking at the donor imputation stages (QRF fits, correlated rank draws, regime gating, NEED raking), the row-stacking and mass-split arithmetic, and the signed-register semantics. Essentially everything found is in the proof and acceptance machinery rather than in the imputation chain — the construction itself held up under scrutiny, and the diagram's attribution rule (last stage to produce or rewrite) made the false-positive classes easy to avoid. That is the good news and also the substance of the review: the spine looks sound, but several of the receipts and gates that certify it can currently pass, fail, or verify nothing for reasons unrelated to whether the data is right. Findings 1–3 are that category, which is why I would fix them before merge rather than after. On the merge bar: for this stage, per-variable parity with the eFRS incumbent doesn't seem like the right criterion — the incumbent isn't the referee for the orange donor-imputation nodes, and its calibrated weights make weighted comparison uninformative by construction. The bar that an approval here would actually be certifying is that the spine composes as declared and that the proofs are real. The first looks true; the second needs 1–3. |
An armed UK national build crashed before its first stage: _source_pins stored the full Ledger provenance block under the 'ledger_facts' role, but role_pins_digest requires each role to carry exactly sha256 and size_bytes, so every armed run raised ValueError at input_pins_digest. The two halves landed independently — the strict pin contract with Logbook adoption (#666) and the Ledger role pin with the compile-parity wiring — and neither PR branch fires it alone, because only an armed run (--ledger-facts) reaches this path and the licensed runs were held. It reproduces on both #743 and #747 branches and on main. The pin now carries the feed's verified digest and byte size; the richer Ledger identity block already travels in safe_artifacts, source_vintages, and the diagnostics build block, so nothing is lost. This fix belongs upstream in #743 (the run-readiness PR whose runbook documents the armed command); it rides the build branch until then.
|
All seven addressed in Blocking1. Surface-vocabulary mismatch — confirmed, and worse than stated. Every committed entry is scoped to Fixed by bridging rather than by re-scoping the register: a 2. 3. Vacuous E7 receipt — confirmed. It now refuses when Follow-ups4. The two spine defects, same commit
Whole-spine parity re-verified after the register semantics changed: On the merge barAgreed, and it is the bar the PR now argues for explicitly. Per-variable parity with the incumbent is the wrong criterion here — the incumbent isn't the referee for the orange donor-imputation nodes, and its calibrated weights make weighted comparison uninformative by construction. What an approval certifies is that the spine composes as declared and that the proofs are real. Per-step gating of each imputation stage, so a regression is caught where it happens rather than at a whole-spine diff, is filed as #757. |
|
Follow-up verification pass over Eight of the nine hold. Verified sound, not just taken on trust: the The bridge is the exception, and it needs one more turn. 1.
|
The frozen instruments were pinned at the incumbent data package's 1.56.14, which carries uk-data#461: from the 2024-25 FRS release the raw benunit table is no longer ordered by sernum, so every benunit-level variable — benunit_id included — landed on the wrong benefit unit relative to the model's sorted-id entity order. The parity screen compares unweighted nonzero shares, which are invariant under a row permutation, so it cannot see this. is_married is one of the 145 columns it compares and reads a byte-identical 0.256587 on both artifacts while every value sits on a different row. Signing whole-spine parity against 1.56.14 would have frozen the upstream defect into the contract as though it were correct. Verified on the new artifact before trusting it: entity counts and column surface unchanged (113,617 / 61,223 / 52,846; 145 layers), clone_index uniformly 0 so it is still the pre-clone artifact that compares row-for-row with the spine grain, and benunit_id now sorted ascending where 1.56.14 was not. The identity is corroborated three ways — the HF LFS metadata at the tagged revision, the repo's own releases/1.56.16/release_manifest.json, and a local hash. Reference-side movement is small: 39 of 145 shares move, none by more than 0.0046, so the +/-0.02 screen is undisturbed. The licensed weighted register moves on all 128 comparable columns, but that is the already-signed register-realization class — 1.56.15 changed the UC caseload targets and the incumbent's calibration re-solves with unseeded dropout — and stays far inside the gross-mass fence. No gate threshold moves here; #686 arms the re-measured baselines later. uk/frs_release.json is deliberately untouched: its revision names the raw UKDS zip, a different artifact that did not change. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
FRS 2024-25 retired CWATAMT and CSEWAMT ("Final value after discount").
The headers survive but carry no data in any of the 16,288 households,
and the FRS replaced them with the Scotland-only CWATAMT1/CSEWAMT1
("Weeklyised gross annual dom. water/sew. charge on bill", "DV created
in 2024-25 as variable was removed from the dataset").
Two things were live. The incumbent adds the retired cell before
filling, so a wholly blank CSEWAMT propagates NaN and zeroes the charge
for all 1,663 Scottish households that have one; our per-column fill
left it standing, which is the whole +0.1009 share divergence flagged
on #736 — we are correct, the incumbent is defective, and the gap
reproduces on the raw tab at +0.1021 before composition. Separately our
own level was short: CWATAMTD is the water charge alone, about GBP 185
per Scottish household against roughly GBP 490 for England and Wales.
Both consumers now call one helper, so the amount netted from the
council tax bill is exactly the amount charged as water and sewerage —
an inconsistency the incumbent still has, since its netting fills per
column while its charge fills after the addition. The helper adds
CSEWAMT1 discounted at the household's own observed CWATAMTD/CWATAMT1,
which keeps the retired cells' after-discount meaning instead of
silently switching to a gross basis. The factor is well behaved: range
(1/3, 1], never above 1, for the 1,641 households with a positive gross
bill. The other two domains cannot be moved by it — the 22 with a
recorded CWATAMTD but no gross cell have zero sewerage, and the 21 with
no council-tax cells stay at zero.
The level fix does not move the nonzero share (0.8783767190569745
either way, matching the existing spine evidence to every digit), so
the share difference stands alone as an incumbent-defect divergence.
The fixtures supplied a non-missing CSEWAMT and so never exercised what
the tab contains; they now carry the retired cells blank, and a
regression test pins that a blank retired cell and an absent one give
the same non-zero answer.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Disclosure-controlled evidence for the two changes above: the pin identity and its three-way corroboration, the structural verification of the new artifact, the defect footprint across benunit/person/ household, the reference-side share movement, the licensed weighted register's digests and relative drift, and the FRS 2024-25 water cell retirement with the factor domain and level comparison. Digests, column counts, unweighted shares and relative deltas only per CD171 5.2.1; the licensed register itself stays uncommitted under the UKDS EUL. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The whole-spine comparison has one rule: anything differing between the spine and the frozen incumbent that is not signed is a defect. That rule needs somewhere to read the signatures from, so this adds the register, its loader, and the validation that keeps it trustworthy. Each entry names its class, the exact surface and columns the difference is expected to appear on, disclosure-safe magnitude evidence, the adjudicator and the date. The loader enforces the three vocabularies, unique kebab-case ids and ISO dates, and a test refuses a column-surface entry that names no columns — an unscoped entry would quietly absorb unrelated divergences, which is precisely the failure this register exists to prevent. It sits above the per-gate reviewed-exclusion registers rather than replacing them. Those are per-gate, per-reference and expiring, because a suppression must not outlive its reason; a signed difference is a permanent adjudicated fact. So expires_on is rejected outright, with the error pointing at the exclusion registers instead. Seeded with the two Scottish water adjudications, deliberately scoped apart: the incumbent's NaN-zeroing signs the share surface, the successor-cell level change signs weighted totals and leaves the share — which it does not move — unsigned. The E4-E8 method classes and the #723 beyond-band columns are transcribed as each is re-measured against the re-pinned reference and scoped to what is actually observed, rather than carried across on prose classification. The country package declares everything it ships, so the resource is registered there; that moves the UK spec bundle sha, re-pinned in the same change. Wheel-checked: the resource ships beside its siblings. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The swap decision rests on this comparison, so both halves of it live here: a candidate mode on the existing extractor, and the diff that holds its output to the signed register. The candidate mode reuses build_reference and build_weighted_totals unchanged and only swaps the identity block. That is the point — if the two sides were measured by different producers, the diff would be between two measurement methods rather than two artifacts. A candidate is identified by its own sha256 instead of being checked against the incumbent pin, and the mode cannot write the committed reference: it refuses any destination inside the country package, and --check is refused outright because a candidate can never satisfy a check against the incumbent pin. verify_uk_spine_parity.py compares the record-count identity exactly, per-column nonzero shares at the reference's own six-decimal grain with the column-set difference both ways, and optionally the two licensed weighted registers as relative deltas only. Every difference must match a register entry; anything else is a defect and exits 1. Two fences stop the verdict being manufactured. The reference side is always the committed instrument, and a candidate extraction claiming the incumbent's own sha256 is refused rather than compared — a copied reference would pass by construction. --strict, the swap-acceptance posture, also fails when a register entry matched nothing, so the register cannot decay into a blanket amnesty as the spine changes. Verified end to end against the E8 spine artifact: 142 columns compared, the household record-count identity holding while the known donor- composition deltas show on persons and benefit units, and water_and_sewerage_charges reproducing at +0.100897 and binding to its signed entry rather than counting as unsigned. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The swap comparison asks a different question from payload identity: the control and candidate builds should share a surface exactly and differ in values only where a difference has been adjudicated. This re-verdicts the same measurements rather than relaxing them. Every structural predicate stays strict — same keys and stored kinds, row counts, column lists in order, dtypes, indexes, root-attribute names — and each differing column or root attribute must name an entry in the committed signed-differences register. A signature excuses a differing value, never a differing surface, which is pinned by a test that supplies a signature for an added column and still expects failure. payload_identical is still computed and reported in both modes, so a structure-only receipt stays comparable with a full-mode one, and --signed-differences is refused outside the mode rather than silently ignored. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Not an acceptance run — the E8 spine predates both the re-pin and the water fix, and the register is deliberately seeded with only the two adjudications made so far. It records that the instrument behaves as specified on real artifacts before the licensed ladder depends on it: the household record-count identity holds, the known donor-composition deltas surface on persons and benefit units, the E9 derived-benefit columns are still the three missing on the candidate side, the water column binds to its signature instead of counting as unsigned, and the weighted-totals entry correctly reports as unused when no weighted registers were supplied. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Both UK resource tuples in test_country_spec.py enumerate the package's resources exactly, so #686's signed-differences register has to appear in them. Caught by the full build-shard run rather than the targeted ones. The first test's name records the #717 question it was written to answer, but what it does now is pin the whole legacy-JSON list; a note says so, so the next reader is not puzzled by a resource arriving in a test that says nothing was added. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
L0 ladder built from raw licensed tabs on this branch: smoke, dev and full all clean, every attempt landing a Logbook row. The record-count identity closes exactly at full scale — (16,288 + 10,000) x 2 + 270 = 52,846 — and the Scottish water fix reproduces at every rung. Parity against the re-pinned reference finds 26 columns beyond the band, against 27 in the #723 screen, which is what a re-pin that moved no reference share by more than 0.0046 should produce. Includes a correction to how divergences are attributed. The surviving value belongs to the last stage to produce *or rewrite* a column, and attributing by producing stage alone manufactures false findings: savings_interest_income and tax_free_savings_income originate in frs_spine and are rewritten by the SPI channel, so a naive pass reports them as raw-mapping divergences outside every signed class — the exact signature that made the water defect real. With rewrites folded in, every beyond-band divergence lands in an established class and the only raw-mapping one is already signed. The queue itself is left unadjudicated. The verdict is defect by construction because the register holds only the two water entries, which is the correct starting state; each class needs a ruling before it becomes an entry, and the E9 derived-benefit gap is the one item that may argue against swapping rather than for signing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Twin determinism passes: two independent full builds are payload-identical across every store key, differing only in bytes because HDFStore stamps write times. All four identity receipts pass the property they exist to prove — that recomputation is invariant to row order. e4, e5 and e6 nevertheless report matches_stored_columns false, and the cause is the instrument, not the spine. _frs_only_frame scopes the survey channel by excluding household_is_spi_synthetic alone, which was right when the SPI channel was the only layer stacking rows. E8 added two more: the capital-gains incidence clone and the 270 band donors. Those rows carry values copied from their sources, so recomputing an identity-keyed draw for a clone's own household id disagrees with a stored value that was never drawn for that id. The measurement is unambiguous: excluding all three flags leaves exactly 16,288 households, the raw FRS count, which is the scope the #723 receipts ran at and passed. e8's own receipt passes because it recomputes the clone and donor logic explicitly instead of assuming unstacked rows, and e4's mismatch list is entirely identity-keyed draw columns — precisely those a clone inherits rather than draws. Recorded as unsigned and unfixed: the scope must exclude every stacked layer, the fix belongs with the e7 receipt work since both are ladder maintenance, and until then these three results say nothing about the spine and the L1 leg of the gate is not satisfied. Also noted: e5 and e6 report the failure with an empty mismatch map, which is not actionable evidence and should name what disagreed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Building a spine that carries every stage through E8 and running the whole ladder surfaced that e4, e5 and e6 had gone stale. Each was written when the SPI support channel was the only stage stacking rows; E8 added the capital-gains clone and the CGT band donors after them. The failures are in the instrument, not the spine, and the experiment that proves it is the unchanged tool against two artifacts: the pre-E8 spine passes e6, the post-E8 spine fails it. Three mechanisms, which is why one fix did not cover them. e4 recomputes identity-keyed draws and a stacked row carries a value copied from its source, never drawn for its own id. e5's regional uprating scales to a per-region mean over the frame's owner households, so stacked rows move the denominator — confirmed by running it both ways on one artifact. e6 fails scoped as well as unscoped: its NHS allocation normalizes against an absolute budget, so it needs stage-time weights, and it divided out the SPI channel's share while the clone's mass_split went unrestored. Scoping now excludes every stacked layer from one declared flag list, and the weight divisor reads its factors from the declared operations instead of hardcoding them. The divisor is driven by the flags the artifact actually carries rather than by the committed roster. The first implementation used the roster and divided the clone factor out of a spine built before that stage existed, skewing the comparison the other way; the pre-E8 artifact caught it at once. Both vintages now pass. Recorded at the helper: a new mass-redistributing op kind has to be registered there, and the failure mode if it is not is a receipt silently comparing against the wrong grossing scale. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two gaps the whole-spine parity screen surfaced, both in already-merged increments rather than in new work. The spine never mapped healthy_start_vouchers, free_school_breakfasts, free_school_fruit_veg or free_school_meals, so parity reported three of them as columns the candidate does not produce. The incumbent maps them straight off the person tapes in create_frs, which makes this an E2/E3 omission rather than deferred work — #685 is UC deduction attributes, bus fares and WAS debt, none of which touch these. heartval is on both the adult and child tapes and the three school columns are child-only, so adults read zero rather than propagating NaN, which the test pins. Measured against the reference the three contract columns land at -0.000089, -0.000001 and +0.000008; free_school_breakfasts is not an engine-known variable so it never enters the 145-column surface. The identity ladder had no e7 check at all — the increment that introduced row stacking, and so the one whose interaction with E8 broke e4, e5 and e6, was the only one without a receipt. It now receipts the support-channel layer: each entity's channel and clone index, the composite source key, and the propagation of a household's channel to its persons and benefit units. Bitwise on both surfaces, since these are labels and integer indices. A negative control confirms it is not vacuous: flipping one channel label fails it and names the column. It deliberately excludes the employer_pension_contributions = 3 x employee_pension_contributions derive. That is a real E7 layer, but E8's salary_sacrifice rewrites the multiplicand in place afterwards and the relation survives on only 95.9% of survey-channel persons, so asserting it would fail for the wrong reason — the #721 rewrites-provenance class. The coverage manifest is regenerated for the new source-manifest hash (145 required, 0 exclusions, unchanged) and the spec bundle sha re-pinned. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The living adjudication packet is a Claude artifact, which is private by default — a reviewer clicking through from the PR hits an access wall. The packet therefore also lives in the repo, where GitHub renders it in the review itself and it travels with the branch. Same content as the interactive page: what the shares measure and their limits, the three incumbents, per-class share tables with donor or raw-source truth, the levels table with the taxable-to-taxable SPI comparison, coverage, diagnostics, and the open items. Adds the UKDS EUL clause 11-12 citations and the disclosure-control statement, which apply wherever these aggregates are posted. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
UC had no dedicated row despite being what the armed calibration binds. Measured modeled universal_credit through the engine on both artifacts: unweighted the spine carries 9% more UC-positive benunits than the incumbent at an equal per-recipient level, and the reported column is exact against the raw tab to every digit. The halved weighted caseload is the calibration boundary itself, not a spine defect — the incumbent's weights already embody a UC caseload target since 1.56.15, so the comparison is before-medicine to after-medicine. Recorded watch-item: would_claim_uc frozen at 0.55 (U8, uk-data#452), the pre-registered lever if the armed run's UC fit is strained. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Main moved uk/target_references.json and uk/target_reference_membership.json, both of which the UK spec bundle hashes, so the pinned bundle digest in the country-bundle test moves with them. Re-cut against the rebased base; the contract mirrors, the parity reference and the coverage manifest all verify unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Takes the whole-spine parity verdict from defect to signed_parity with nothing unsigned and no strict failure. Two changes make that honest rather than merely green. The register grows to thirteen entries, split so that no entry covers both columns where the spine is closer to its donor and columns where the incumbent is. Donor evidence is re-measured through each stage's own committed cleaning function on the survey-weighted basis, which reproduces the E6 acceptance receipt's education figure exactly; that settles the ETB weight-basis question and exposes the incumbent's dfe_education_spending as degenerate at fourteen nonzero households in 52,846. The instrument's share surface moves from the reference's six-decimal grain to the #723 acceptance band. Ninety columns sit inside it on third-decimal drift, and signing those would have been the blanket amnesty the register is built to prevent. Nothing is hidden: in-band differences are reported under their own key with the in-band maximum, --share-band 0 restores the exact check, and structural differences stay outside the band's reach at any band. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Re-measures every donor comparison through each stage's own committed cleaning function on the survey-weighted basis, adds the level ratios alongside the shares, and replaces the E6 "needs ruling" and E5 "carry-forward?" sections with what the evidence now shows. The ETB rows are no longer undecidable: the stage cleans a thirteen-column subset of one year and weights by hhold_adj_weight, and on that frame the incumbent's education column is degenerate at fourteen nonzero households in 52,846. The incumbent's per-head division is recorded as a second upstream defect, observed but not filed. Also records the correction that a mid-review unweighted re-measurement of the wealth columns appeared to overturn the standing E5 adjudication and did not: WAS oversamples wealth-holders, so the weighted basis is the population one, and on it the original reading stands. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
R5 receipts the signed_parity verdict and, more usefully, the two ways the donor measurements had been wrong before it. Both had the same cause: a donor truth was computed on a frame the stage does not use. Once against the ETB services frame, which produced a false blocker and had been recorded as an unpinnable weight basis; once against the WAS frame unweighted, which briefly appeared to overturn a standing adjudication. Calling each stage's own cleaning function reproduces the E6 acceptance receipt's education figure to four decimals, which is the check that the convention is the right one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The signing tables report population mean per household; the Levels section reports conditional mean per carrier. They can disagree on which side is closer, so the section now says which one the signing used and why: the population mean is what a calibration target binds and what survives a difference in incidence. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
An earlier pass read the only full incumbent H5 on disk, which is the 1.56.14 published artifact rather than the pinned 1.56.16 one. The two differ by at most 0.0045 on any share, but that was enough to flip one verdict: alcohol_and_tobacco_consumption reads as ours at 1.56.14 and as the incumbent at the pin. It moves out of the donor-faithful entry into a two-column entry with transport_consumption, which shares its evidence shape exactly — incumbent closer on share, ours closer on level. The pinned artifact is now fetched and verified by digest before measuring, and the correction is recorded in the ledger rather than quietly applied. Also rewrites "What the numbers are", which still claimed the ledger measures incidence and never level. That is true of the committed instrument and is precisely why the ledger exists; the document itself evaluates both, and the section now says which quantity is operative when they disagree. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A cross-check of every register-quoted number against its measured source found the incumbent side had been measured on the 1.56.14 artifact staged during #723 rather than the pinned 1.56.16 one. The instrument itself was never wrong, since it reads the committed reference; the hand measurements beside it were. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The earlier pin correction updated the level ratios and the other-residential share but left savings, property_wealth and corporate_wealth quoting the 1.56.14 incumbent. Caught by re-running the cross-check that compares every register-quoted number against its source. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Eleven of the thirteen evidence fragments did not match any heading in the file they cite — two of them since before this branch, the rest because the ledger's section titles changed when the queue was signed. A dead fragment fails silently in a browser, so the rot only surfaces when a reviewer clicks and lands nowhere, which is exactly when the pointer needed to work. The new test derives GitHub's own anchor form from each cited file's headings and requires the fragment to be among them, so renaming a section now breaks CI instead of breaking an audit trail. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two lanes of fixes that belong together: the review found the spine's
construction held up and its proof machinery did not, and the first armed
calibration campaign found two spine defects it had to work around outside
the build. Both are cheaper to fix now than after a swap.
Spine defects. The SPI stage left twelve full-concept income columns NaN on
the FRS channel, which made the artifact unloadable in practice — the engine
refuses NaN inputs and any finiteness fence on the calibration path refuses
the frame — so the campaign zero-filled them outside the build. Zero is the
adjudicated stage-time semantics for a concept the instrument never asked
about. Separately, the licensed FRS records no age above 80, so the 85+
population targets were structurally unbindable and the 80-84 band carried
the whole 80+ population; the new age_tail stage disperses the pile from a
sex-specific inverse CDF over committed ONS band populations, keyed on
person_source_id so clone twins agree, running last so nothing that
conditions on age sees a different input.
Proof machinery. Three findings could change what a proof means: the
register and the payload comparator spoke different surface vocabularies, so
--structure-only could only ever fail; `expectation` was validated and then
never consulted, so a column's signed *appearance* also signed its values;
and the E7 receipt reported green over an empty recomputation. Four quieter
ones: a candidate omitting its source identity passed the aliasing fence
vacuously, a zero reference total produced float("inf") that the encoder
refused, --strict false-failed on within-band signed columns, and the
Scottish water discount fallback rested on an unasserted vintage claim.
Whole-spine parity re-verified after the register semantics changed:
signed_parity, 0 unsigned, --strict clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CI caught eight failures I should have: adding age_tail moves three derived surfaces, and I ran targeted suites instead of the shard. The stage roster is pinned in three places, and the E8 contiguity invariant asserted that the E8 block sat immediately before the certified pair — true only because E8 happened to be the last increment. What the invariant is actually protecting is that E8 stays contiguous and the certified pair stays at [-2:], where the frozen-copy lockstep test reads it; both survive a spine tail. The test now says that, with age_tail declared as a POST_E8 block so the next stage after it has to make the same decision deliberately. The release input-coverage manifest carries the source manifest's digest, so it regenerates: 145 required columns and 0 exclusions unchanged, confirming age_tail adds no column — it rewrites `age`, which was already required. The stale digest was also what failed the preflight battery's manifest-current gate. Verified against the packaging gate this time: wheels built, installed into a clean venv, suite run from there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…lted The re-review found the payload bridge made an old gap load-bearing: covers() matched on surface and column only, so a household-scoped adjudication could sign a same-named column on the person or benunit table once a per-table comparator started consulting it. That is an unsigned divergence becoming silently signed — the failure the register exists to prevent, reintroduced by the fix for it. Lookups now carry the entity, the payload comparator passes the store key it is iterating, and the bridge forwards the entity scope rather than widening it. The check earned its keep immediately: student_loan_balance was scoped to household and is a person column. The loader now also refuses a column-surface entry that names no columns or no entities, so the surface-wide form survives only on entity_counts, where the "column" is itself an entity name. Weighted-totals matches were already counted in matched_ids before unused is computed; a test now pins that a totals-scoped entry reads as matched rather than as rot, since the surface is dormant and the accounting would otherwise first be exercised on a calibrated candidate. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
86f5574 to
76e39f9
Compare
|
Fixed in 1.
|
vahid-ahmadi
left a comment
There was a problem hiding this comment.
Approving. 76e39f9b is green on all four checks and is the commit I reviewed.
I verified the entity-scoping fix in the diff rather than reading the write-up: covers() and matching() take an entity, and every call site now passes one — the payload comparator passes its store key, _compare_shares and _missing pass entities.get(column), and _compare_weighted_totals gained an entities parameter fed from reference.input_entities. The bridge forwards the scope rather than widening it. Refusing an unscoped column-surface entry at the loader is a stronger fix than the one I suggested, and entity_counts is correctly exempt since its "column" is the entity. The re-scoped was-student-loan-balance-fold entry has its magnitude_evidence prose updated to match rather than left stale, and the test asserting every committed entry's entities agree with the reference's input_entities is what stops the next one sitting latent.
Your push-back on finding 3 was correct and the finding was wrong. I read verify_uk_spine_parity.py at 76e39f9b: matched_ids.update(...) from totals_report["differing"] is inside the optional-totals block and unused is computed after it, so a totals-scoped entry that matched has always counted as matched. The pinning test is a good outcome from a bad finding.
That the entity check immediately caught student_loan_balance scoped to household when it is a person column is the substantive result here — it means the previous 0 unsigned was reachable with a real mis-scope in the register, and the current one is not.
Worth stating what this approval covers, since the PR is careful about that distinction elsewhere: I reviewed the diff and confirmed CI is green. I did not run the build, the parity instrument, or the suites locally, so the signed_parity / 0 unsigned / --strict clean result is one I have read the code for rather than reproduced. On that basis the spine composes as declared and the proofs now mean what they say.
…the reformatted manifest and re-cut the digests over the union Main's #747 reformatted uk/source_stages.json and moved the attested surfaces again, so the merge takes main's file and re-applies the five WAS-stage edits (derived split, chain order, outputs, nonnegative outputs, notes incl. the per-segment seed sentence) in its format, regenerates release_input_coverage_manifest.json (145 required / 0 exclusions unchanged), re-pins the UK spec_sha256 and re-cuts the three gate-battery digests over the union - the d70ea39 pattern, second application. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Whole-spine parity instruments and evidence for #686 (WS-E E10, epic #665). This PR delivers the comparison layer that decides whether the microcosm-built spine can replace the incumbent, plus two fixes the comparison itself surfaced.
Instruments
enhanced_frs_2024_25.h5carried uk-data#461 (raw benunit table no longer sernum-ordered → every benunit-level variable misassigned). The parity screen compares nonzero shares, which are permutation-invariant, so it is structurally blind to this defect —is_marriedreads a byte-identical 0.256587 on both artifacts while every value sits on the wrong benefit unit. Re-pinned to 1.56.16 (adjudicated; identity corroborated three ways), all instruments and the four gate-battery mirrors re-cut. Reference-side movement ≤0.0046 per column, so the ±0.02 screen is undisturbed. The spine itself was never exposed:frs_spinehas sorted the benunit table by id since its original ingest commit.CWATAMT/CSEWAMT; the incumbent's fill-after-addition NaN-zeroes the charge for all 1,663 Scottish households (the whole +0.1009 divergence — incumbent-side, reported upstream as policyengine-uk-data#467). Our own level was also short (CWATAMTDis water-only); both the spine mapping and the council-tax netting now share one helper addingCSEWAMT1at the household's own observed discount factor. Verified at every rung: share unchanged at 0.878377, Scottish level ~£395/household vs ~£185 before.uk/spine_swap_signed_differences.json+ loader): the migration-scoped record of adjudicated spine-vs-incumbent differences, seeded with the two water entries. Strict vocabulary, unique ids, precise per-surface scoping, no expiry (permanent adjudications; expiring suppressions stay in the per-gate exclusion registers). Advisory, not a gate — the standing health checks are the existing battery (support bounds, tail concentration, input-mass parity, aggregate-admin anchors, take-up signal, enum domains), per the US precedent; this register documents the one-time swap decision.tools/verify_uk_spine_parity.py— diffs a candidate extraction against the committed reference on entity counts, nonzero shares, and (optionally, licensed) weighted totals; every difference must match a register entry or the verdict isdefect. Fenced against manufactured verdicts: the reference side is never derived from the candidate, a candidate claiming the incumbent's sha is refused, and--strictfails unused register entries.build_uk_efrs_parity_reference.py --candidate-h5— extracts the same-shape surface from a candidate spine with the same producer (same engine, aliases, rounding), structurally unable to write the committed reference.compare_uk_h5_payload.py --structure-only— additive verdict mode for the eventual control-vs-candidate comparison: structural predicates stay strict, value differences must be signed.--check e7receipts the support-channel layer (the one stacking increment that had no receipt); a negative control proves it non-vacuous. e4/e5/e6 had gone stale against E8's stacking — three distinct mechanisms (identity-keyed draws on copied rows; a population-dependent regional-uprating mean; the NHS budget normalization needing stage-time weights) — fixed with scoping driven by the artifact's own stacking flags and mass factors read from the declared operations. Green on both the post-E8 and pre-E8 spines.free_school_meals,free_school_fruit_veg,healthy_start_vouchers,free_school_breakfasts): a raw-mapping omission from the merged E2/E3 port, previously misattributed to E9. Coverage is now 145/145; the three contract columns land within 0.0001 of the reference.Evidence (licensed,
data/ukds/acceptance/686-spine-swap/; receipts inexperiments/686-uk-spine-swap-receipts.md)rewrites, not justproduces), with the only raw-mapping divergence (water) already signed.Comparison ledger (adjudication packet): committed in this branch as
experiments/686-uk-spine-comparison-ledger.md— GitHub renders it directly. An interactive, continuously updated version lives as a Claude artifact (private by default; ask María for access). The committed rendition carries the UKDS EUL clause 11–12 citations.The queue is signed
verify_uk_spine_parity --strictnow returnssigned_parity: 26 beyond-band divergences, all signed; 0 unsigned; no strict failure; household count exact.The register grows from 2 entries to 13, and how they are split is the substance. No entry covers both columns where the spine is closer to its donor and columns where the incumbent is — so the LCFS class is three entries, not one: nine columns where the regime-gated draw lands closer to the donor,
petrol_spending/diesel_spendingwhere thehas_fuelgate under-places incidence and the entry says so, andtransport_consumption/alcohol_and_tobacco_consumptionwhere the incumbent is closer on share and the spine on level. The four incumbent-closer columns each carry that direction in their own entry text, so re-opening one does not require re-opening the class.Three things the re-measurement changed.
The ETB rows were never undecidable. The weight basis was not the problem — the frame was. The services stage cleans one year on complete cases over a 13-column subset (4,199 rows); the incumbent cleans an 18-column subset, so a "donor share" off its frame has a different denominator. On the stage's own frame the donor share is 0.2546 unweighted — the acceptance receipt's figure to four decimals. With that fixed, the incumbent's
dfe_education_spendingis degenerate: 14 nonzero households in 52,846, £2 per household against a donor £3,461. Signed as a defect fix on the incumbent side.A reversal that was itself wrong. An unweighted re-measurement of the wealth columns briefly appeared to overturn the standing 2026-08-19 E5 adjudication. It did not — WAS oversamples wealth-holders by design, so its unweighted frame is not a population. On the weighted basis the original reading reproduces exactly, and
corporate_wealthgains a benchmark it previously lacked. Both near-misses had one cause, recorded as a standing lesson in R5: a donor truth is a property of the frame and weighting the stage itself uses, not of the donor file.The incumbent side was measured on the wrong artifact, caught by a cross-check of every register-quoted number against its measured source. The only full incumbent H5 on disk is the 1.56.14 one staged during Retarget the UK build to FRS 2024-25: re-pin raw vintages, regenerate parity instruments, re-measure gate baselines #723, not the pinned 1.56.16 one; the pinned artifact is now fetched and digest-verified before measuring. The two differ by at most 0.0045 on any share — enough to flip
alcohol_and_tobacco_consumption, which moved out of the donor-faithful entry into a two-column entry withtransport_consumptionthat shares its evidence shape exactly. Level ratios moved in the third digit; no other verdict changed. The instrument was never wrong here — it reads the committed reference, which was always at the correct pin — the hand measurements beside it were.One instrument change reviewers should look at
The share surface is now held to the ±0.02 #723 acceptance band rather than the reference's 6-decimal grain. Ninety further columns sit inside it on third-decimal drift from re-running every stochastic stage; signing those would have been exactly the blanket amnesty the
--strictfence exists to prevent — 90 permanent adjudications describing noise, each blanket-covering its column against any future real regression.The band governs only which magnitudes require an adjudication. Every difference down to the 6-decimal grain is still in the receipt under
within_bandwith the in-band maximum (0.019) alongside;--share-band 0restores the exact check; and structural differences are outside the band's reach at any band — a column appearing or vanishing, and every entity count, still signs exactly. There is a test for that specifically, run at--share-band 0.9.--strictalso now separates an entry that matched nothing on a surface this run compared from one whose surface was never examined. The weighted-totals surface stays unexamined until there is a calibrated candidate — comparing our design weights to the incumbent's calibrated ones would flag all 131 columns, the same before-medicine/after-medicine error the UC check documents — so the water-level entry reports as dormant, and a test pins that dormancy stops being available the moment a run supplies the sidecars.Open, and deliberately not decided here
communication_consumptionis the one LCFS column where the incumbent is closer on level while we are closer on share. Signed inside the donor-faithful class; scoping it out is a one-line register change.student_loan_balancehas no like-for-like donor benchmark — signed on the standing E5 adjudication with its direction unevidenced, stated rather than papered over.Rebased, and what follows
Rebased onto
main(239 commits). The only derived value that moved was the UK spec bundle digest, because main toucheduk/target_references.jsonanduk/target_reference_membership.json.Spine-lane follow-up is now recorded on #757 (rescoped 2026-08-25 against the #623 calibration-seam plan): the spine defect batch the first armed calibration campaign surfaced (structural-NaN SPI columns — which the seam's NaN fence will refuse until fixed — the 85+ age-tail source stage, sidecar self-pin, support-channel gaps), per-step gating of the spine build per the US #327 doctrine with the
required_at_buildre-point, and the W9–W11 retirement with the markdown cleanup and the candidate-name move tomicrocosm_uk_2024. The calibration-only driver itself moved out to #623 (PR #743): the new seam calibrates whichever h5 it is given and never modifies a data variable, and the legacy June path retires here once the seam is proven. #757 is explicitly gated on this PR being reviewed and merged.Refs #686, #736, #145, #665. Upstream: policyengine-uk-data#467.
🤖 Generated with Claude Code